Papers with dialectal Arabic
Arabic Diacritization Using Morphologically Informed Character-Level Model (2024.lrec-main)
Copied to clipboard
Muhammad Morsy Elmallah, Mahmoud Reda, Kareem Darwish, Abdelrahman El-Sheikh, Ashraf Hatim Elneima, Murtadha Aljubran, Nouf Alsaeed, Reem Mohammed, Mohamed Al-Badrashiny
| Challenge: | Diacritics are typically omitted in Arabic writings and the reader needs to guess the proper diacritics as they are reading. |
| Approach: | They propose a morphologically informed character-level model that can recover both types of diacritics simultaneously. |
| Outcome: | The proposed model achieves lowest word-level diacritization error rate for Classical Arabic, MSA, and two dialectal Arabic texts. |
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic (2024.acl-long)
Copied to clipboard
| Challenge: | Existing algorithms for recognizing borrowed and dialectal sounds are limited to Arabic, a dialect-rich language containing more than 22 major dialects. |
| Approach: | They propose a framework to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages that extends beyond its standard orthographic sound sets. |
| Outcome: | The proposed framework improves character error rate by 7% with only one and half hours of training data compared to the baseline. |